Papers with end-to-end accuracy
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models (2024.emnlp-main)
Copied to clipboard
Yash Akhauri, Ahmed AbouElhamayed, Jordan Dotzel, Zhiru Zhang, Alexander Rush, Safeen Huda, Mohamed Abdelfattah
| Challenge: | Prior work has focused on contextual sparsity, but it has not been successful. |
| Approach: | They propose a novel pruning predictor that can shadow the LLM behavior and enforce better sparsity patterns. |
| Outcome: | The proposed model can shadow the LLM behavior and enforce better sparsity patterns, resulting in 15% improvement in end-to-end accuracy compared to prior methods. |
Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent work emphasizes improving efficiency in LLM-based systems, especially for longcontext and multi-step reasoning. |
| Approach: | They analyze the role of listwise reranking in deep search pipelines and compare their results to a novel ETC metric to determine model scale and reasoning effort. |
| Outcome: | The proposed model scale, reasoning effort, reranking depth, and total token cost (ETC) metric improve retrieval and end-to-end accuracy and moderate reranked agents achieve comparable accuracy at substantially lower cost. |